Papers by Jin Peng Zhou
Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluations (2025.emnlp-main)
Copied to clipboard
| Challenge: | a lack of trust in graders on graduate-level physics and Olympiad-level math makes them unreliable grader. |
| Approach: | They propose to use a grader LM to evaluate the candidate LMs. |
| Outcome: | The proposed approach outperforms human graders on *RewardBench* and human expert grader on Olympiad-level math problems. |